Skip to content

[ROCm][MoE] Pad the AITER MoE intermediate size at allocation time, and round the expert-group count to a kernel that exists - #55368

Merged
shen-shanshan merged 4 commits into
vllm-project:mainfrom
sshlyapn:sshliapn/features/k2_horizon_fixes
Sep 30, 2026
Merged

shen-shanshan merged 4 commits into
vllm-project:mainfrom
sshlyapn:sshliapn/features/k2_horizon_fixes

Conversation

@sshlyapn

@sshlyapn sshlyapn commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Two independent startup failures on the ROCm AITER MoE path, both caused by shapes AITER has no compiled kernel for, both focusing on IFM/K2-Horizon-375B-A23B model support.

1. Unaligned per-partition intermediate size - unquantized path only, fp8 AITER alignment is unchanged

AITER's CK 2stages MoE kernel dispatches on inter_dim <= 192: below it both stages use 64-wide tiles, above it at least one stage uses a 128-wide tile. CK's IsSupportedArgument rejects an intermediate not divisible by that width, so moe_intermediate_size = 1792 at TP=8 gives 224 and fails with device_gemm ... does not support this GEMM problem.

Fix: round the intermediate up in maybe_roundup_sizes, the existing hook for this (already overridden by five other quant methods). It runs before create_weights, so the padded size reaches allocation and the loader — which derives shard offsets from the checkpoint, not the parameter shape — fills only the real rows.

Alignment rule: 64 if inter_dim <= 192 else 128, mirroring AITER's own threshold rather than always using 128. A flat 128 mispredicts 2 of the 19 sizes measured against the real kernel (64 and 192 both pass despite not being 128-aligned) — it is a wrong model that happens to be safe. It also costs: inter_dim is a tuned-config lookup key, and across the 148 shipped AITER config files (11279 rows) the threshold rule pads 0 tuned rows while a flat 128 would move 207 off theirs onto the heuristic fallback. Rounding up never crosses the boundary (96->128, 160->192, 192->192 all stay <= 192), so the alignment picked is always the one the dispatcher then applies. Dtype-independent: CK derives both tile widths from sizeof(A0DataType), identical for fp16 and bf16.

Zero-init. The loader narrows to the checkpoint extent and never writes the tail, so w13/w2 are allocated with torch.zeros when either dim is padded — unpadded layers keep torch.empty and pay nothing. Pad lanes are then bit-exactly inert: silu(0) * 0 = 0, and the zero columns in w2 contribute exactly zero to stage 2. With torch.empty the tail holds recycled device memory, which is live weight to the kernel; whether that hurts depends on load-time allocation order, so it presents as sporadic implausible tokens with no error. mxfp4 and quark nvfp4 already allocate zeroed; #55251 is doing the same for fp8.

Scope. Gated on UnquantizedMoeBackend.AITER, only selected on ROCm with AITER MoE enabled; every other unquantized backend falls through the base hook untouched. The quantized AITER paths have their own overrides with different constants and are not touched — fp8 needs 256.

Known interaction. Since intermediate_size_per_partition now differs from ..._unpadded, experts/rocm_aiter_moe.py computes a non-zero intermediate_pad where it previously got 0. Traced: with QuantType.No + bf16 the dispatch falls past every cktile/flydsl branch (all require q_type == per_1x32) to the plain ck_moe_stage1/ck_moe_stage2_fwd partials, which take no pad argument, and intermediate_pad only enters the tuned-config key when config_file is not None, which vLLM never sets. So it is computed and discarded, and the kernel runs the full padded GEMM including the zero rows — up to ~14% extra compute on the MoE GEMMs (448->512 at TP=4, 224->256 at TP=8).

2. AITER biased_grouped_topk expert-group count

_aiter_get_num_expert_group ceil-divides num_experts by the 32-per-group limit then walks to the next divisor, which can produce a count with no compiled kernel: AITER instantiates biased_grouped_topk only for NUM_GRP in {1,2,4,8}, so 96 experts gives 3 and fails to launch.

Grouping is a no-op here (topk_group == num_expert_group), so any divisor within the limit routes identically — the constraint is purely which kernel exists. Round to the largest supported count that divides num_experts and respects the limit. When none fits (320 -> 10, 33 -> 3) the naive value is kept deliberately: it is large enough that the topk >= num_expert_group guard at the call site fails and routing falls back to the generic path rather than reaching AITER.

Test Plan

Both files are CPU-only and gated on current_platform.is_rocm(), so they run in any ROCm CI job without a working AITER runtime.

tests/kernels/moe/test_rocm_aiter_moe.py — new "Weight alignment" section appended to the existing file rather than a standalone module, following that file's conventions (function-local imports, default_vllm_config, rocm_ naming):

  • test_aiter_moe_alignment_follows_threshold — the rule, as a table across the boundary.
  • test_aiter_moe_padded_size_stays_in_its_dispatch_branch — self-consistency: (intermediate <= 192) == (padded <= 192) and alignment(padded) == alignment(intermediate).
  • test_aiter_moe_roundup_pads_intermediate — the hook end-to-end: 96->128, 192->192, 224->256, 448->512, 512->512.
  • test_aiter_moe_roundup_is_not_applied_to_other_backends — TRITON and FLASHINFER_CUTLASS keep 224.
  • test_aiter_moe_padded_weights_are_zero_initialized — pad lanes are zero, not allocator garbage.
  • test_aiter_moe_aligned_weights_are_zero_initialized — zero-init is unconditional, so it holds for an aligned size too.
  • test_aiter_moe_bias_matches_padded_weight — w13_bias.shape[1] == w13_weight.shape[1] with has_bias=True; the defect the post-load approach had.
  • test_aiter_moe_padding_is_numerically_inert — reproduces the loader's gate/up placement into a zeroed padded parameter, asserts the MoE output is bit-identical to the unpadded reference.

tests/kernels/moe/test_rocm_aiter_num_expert_group.py (new) — five tests over _aiter_get_num_expert_group: the two invariants the router asserts unconditionally, that a supported count is chosen whenever one fits, a table of known expert counts, and that the fallback is taken deliberately when none fits.

Test Result

gfx950 (MI355X), torch 2.11.0+gitd0c8b1f, on 29af8bd672.

pytest tests/kernels/moe/test_rocm_aiter_num_expert_group.py -q   49 passed, 2 skipped
pytest tests/kernels/moe/test_rocm_aiter_moe.py -q                 1 failed, 63 passed
both files                                                         1 failed, 112 passed, 2 skipped  (72s)

The one failure, test_aiter_fused_moe_mi3xx_fp8_accuracy, is pre-existing and unrelated — verified by running it alone against an unmodified checkout of the parent commit in a separate worktree, same box and image, where it fails identically.

@coderabbitai

coderabbitai Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@mergify mergify Bot added the rocm Related to AMD ROCm label Sep 4, 2026
@github-project-automation github-project-automation Bot moved this to Todo in AMD Sep 4, 2026
@sshlyapn
sshlyapn force-pushed the sshliapn/features/k2_horizon_fixes branch 5 times, most recently from eca031f to b619cb1 Compare September 8, 2026 06:41
@sshlyapn
sshlyapn marked this pull request as ready for review September 8, 2026 06:46

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@sshlyapn
sshlyapn force-pushed the sshliapn/features/k2_horizon_fixes branch from b619cb1 to 837ae11 Compare September 8, 2026 06:50
@tjtanaa

tjtanaa commented Sep 8, 2026

Copy link
Copy Markdown
Member

@sshlyapn will following the MoE tuning guide (online tune AITER_ONLINE_TUNE=1 and offline tuning) in AITER repo resolve this issue? I would like to avoid adding overhead by padding unnecessarily if there are kernels that are actually able to handle the size when there are issues.

Line where ONLINE TUNE feature exists.
https://github.com/ROCm/aiter/blob/eec768a47c1da52e175f9089c640e695f8b572a3/aiter/fused_moe.py#L2440

@sshlyapn

sshlyapn commented Sep 8, 2026

Copy link
Copy Markdown
Contributor Author

@sshlyapn will following the MoE tuning guide (online tune AITER_ONLINE_TUNE=1 and offline tuning) in AITER repo resolve this issue? I would like to avoid adding overhead by padding unnecessarily if there are kernels that are actually able to handle the size when there are issues.

Line where ONLINE TUNE feature exists. https://github.com/ROCm/aiter/blob/eec768a47c1da52e175f9089c640e695f8b572a3/aiter/fused_moe.py#L2440

Thank you for comment!

Unfortunately no - tuning can't fix this, because the tile width isn't a tuned parameter. CK instance is selected from inter_dim directly (inter_dim <= 192 then 64-wide, else 128-wide), in both stage-1 and stage-2 dispatch, and CK requires inter_dim to be divisible by whichever N tile it picks, so no tuned config can land inter_dim=224 - IsSupportedArgument rejects it and we get "device_gemm ... does not support this GEMM problem". It's a kernel implementation requirement and it's scoped to the bf16/fp16 data types only.

@sshlyapn

Copy link
Copy Markdown
Contributor Author

Hi @tjtanaa @AndreasKaratzas could you please take a look at this PR when you have a moment?

@simondanielsson simondanielsson left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks similar to #55251, can we adopt the same type of changes?

This also emphasizes that perhaps we should fix this on the aiter side instead. WDYT?

@ChuanLi1101

Copy link
Copy Markdown
Collaborator

Allocation-time pad + zero tails is the right place to fix the CK shape reject, and it avoids the bias-length bug of a post-load rebuild.

Please address before merge:

  1. Rounding num_expert_group to a compiled {1,2,4,8} needs a numerical check vs the Python router (e.g. 96 experts), not only “the kernel exists”. Grouping is a no-op only if topk_group == num_expert_group still selects the same experts.
  2. Call out (or fix) that the unquant CK path still GEMMs the padded intermediate (~14% extra on those GEMMs). Prefer AITER-side pad / the same pattern as [ROCm] Add gelu_tanh to the AITER fp8 fused MoE and zero-allocate the padded expert weights #55251 if we can avoid paying that in vLLM.
  3. This hook is unquantized-AITER only — please note in the PR that fp8 AITER alignment is unchanged.

@sshlyapn

Copy link
Copy Markdown
Contributor Author

@simondanielsson @ChuanLi1101, thanks for the comments!

This looks similar to #55251, can we adopt the same type of changes?

Yeah, the changes are really close, but they address different data types: unquantized BF16/FP16 vs FP8 in the mentioned PR. The logic itself looks the same, I just preferred not to zero out the buffers when the hidden dimension is also padded, as that's not what the MoE triggers/requires - so I'd prefer to keep the scope of the changes narrower

This also emphasizes that perhaps we should fix this on the aiter side instead. WDYT?

Agree, the better fix would definitely be improved unaligned-shape support in AITER. I've created a ticket to track the issue from AITER side: ROCm/aiter#5444. Meanwhile, since that might take some time, I'd suggest keeping this padding to provide functional support for models where such unaligned shapes might appear

Rounding num_expert_group to a compiled {1,2,4,8} needs a numerical check vs the Python router (e.g. 96 experts), not only “the kernel exists”. Grouping is a no-op only if topk_group == num_expert_group still selects the same experts.

Corresponding test has been added, thanks!

Call out (or fix) that the unquant CK path still GEMMs the padded intermediate (~14% extra on those GEMMs). Prefer AITER-side pad / the same pattern

Absolutely agree. Here is the AITER issue for further improvements in this direction: ROCm/aiter#5444

This hook is unquantized-AITER only — please note in the PR that fp8 AITER alignment is unchanged

Done

@AndreasKaratzas AndreasKaratzas added the verified Run pre-commit for new contributors without triggering other tests label Sep 11, 2026
@sshlyapn
sshlyapn force-pushed the sshliapn/features/k2_horizon_fixes branch 2 times, most recently from 99b3b6f to 409c7f7 Compare September 14, 2026 09:19
@mergify

mergify Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @sshlyapn.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@github-actions

Copy link
Copy Markdown

❌ This PR is 97 commits behind upstream main. Your branch must contain every commit currently on upstream main. No new CI build was started. Merge or rebase onto the latest main, then rerun /ci run. To test this branch at your own risk, use /ci run --allow-stale.

@shen-shanshan

Copy link
Copy Markdown
Contributor

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #90117 for commit 2ecf5d309b12.

@shen-shanshan

Copy link
Copy Markdown
Contributor

@sshlyapn The DCO check failed. Please use git commit -sm "...".

@sshlyapn
sshlyapn force-pushed the sshliapn/features/k2_horizon_fixes branch 2 times, most recently from 1ccb361 to 0e2d02f Compare September 21, 2026 05:40
@sshlyapn

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

❌ This PR is 1 commit behind upstream main. Your branch must contain every commit currently on upstream main. No new CI build was started. Merge or rebase onto the latest main, then rerun /ci run. To test this branch at your own risk, use /ci run --allow-stale.

@sshlyapn
sshlyapn force-pushed the sshliapn/features/k2_horizon_fixes branch from 0e2d02f to 2f98b2c Compare September 21, 2026 05:55
@sshlyapn

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #90153 for commit 2f98b2c2f430.

@sshlyapn
sshlyapn force-pushed the sshliapn/features/k2_horizon_fixes branch from 2f98b2c to edf6611 Compare September 21, 2026 08:53
@sshlyapn

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #90186 for commit edf6611aaf2b.

unpadded_hidden = self.moe.hidden_dim_unpadded
assert unpadded_hidden is not None
unpadded_up = unpadded_intermediate * (2 if self.moe.is_act_and_mul else 1)
requires_zero_alloc = (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@sshlyapn can you check if this would cause a problem when we're doing RL reloads? I don't think the pads get zeroed out now when we reload the weights.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@dllehr-amd good catch, thank you! I have checked this and zeroing at allocation didn't survive reload. I have moved it to Aiter branch of convert_to_unquantized_kernel_format before the shuffle, and now it uses the same path as TRT-LLM

AITER's CK 2stages MoE kernel dispatches on inter_dim <= 192: below it both
stages use 64-wide tiles, above it at least one stage uses a 128-wide tile.
CK's IsSupportedArgument rejects an intermediate size not divisible by that
width, so TP splits like 1792/8=224 fail with "device_gemm ... does not
support this GEMM problem".

Round the per-partition intermediate up in maybe_roundup_sizes instead of
rebuilding the weights after loading. The hook runs before create_weights,
so the padded size reaches allocation and the loader (which derives shard
offsets from the checkpoint) fills the real rows.

Fixes two defects in the post-load approach:
  - w13_bias/w2_bias were sized off the unpadded intermediate, leaving the
    bias shorter than the weight it is added to when has_bias is set.
  - layer.intermediate_size_per_partition kept the pre-pad value while
    moe_config was updated, desyncing the two.
It also drops a duplicate allocation of w13/w2 per layer at load.

Allocating padded means the loader never writes the tail, so w13/w2 are
zero-initialized when the dim is padded; the pad lanes must be inert
(silu(0) * 0 = 0, matching the zero columns in w2).

Mirror AITER's threshold rather than always aligning to 128: inter_dim 192
is unaligned but valid, and inter_dim is a tuned-config lookup key. Across
the 148 shipped AITER config files (11279 rows) the rule pads no tuned row;
a flat 128 would move 207 of them off their tuned entry.

Gated on UnquantizedMoeBackend.AITER, only selected on ROCm with AITER MoE
enabled. The rule is dtype-independent because CK derives both tile widths
from sizeof(A0DataType), identical for fp16 and bf16, and the kernels
TORCH_CHECK that the output dtype is one of those two.

Tested on vllm/vllm-openai-rocm:nightly-8a728663c (gfx950):
  pytest tests/kernels/moe/test_rocm_aiter_moe.py  -> 64 passed

Signed-off-by: Sergei Shliapnikov <sergei.shliapnikov@amd.com>
Signed-off-by: Sergei Shliapnikov <sergei.shliapnikov@amd.com>
…ibuteError: 'UnquantizedFusedMoEMethod' object has no attribute 'unquantized_backend'

Signed-off-by: Sergei Shliapnikov <sergei.shliapnikov@amd.com>
@sshlyapn
sshlyapn force-pushed the sshliapn/features/k2_horizon_fixes branch from edf6611 to bd385d1 Compare September 25, 2026 12:22
…oads keep it clean

Signed-off-by: Sergei Shliapnikov <sergei.shliapnikov@amd.com>
@sshlyapn
sshlyapn force-pushed the sshliapn/features/k2_horizon_fixes branch from bd385d1 to 157a3c4 Compare September 25, 2026 12:25
@sshlyapn

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #91202 for commit 157a3c40ba0a.

@sshlyapn

Copy link
Copy Markdown
Contributor Author

/ci retry

@github-actions

Copy link
Copy Markdown

✅ Queued 2 failed job(s) for retry in Buildkite CI #91202.

@shen-shanshan shen-shanshan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM, since it shows performance gain in most scenarios. Btw, could you please add the gsm8k accuracy results before merging?

@sshlyapn

Copy link
Copy Markdown
Contributor Author

@shen-shanshan here are results for K2-Horizon-375B-A23B bf16, TP8, gsm8k 5-shot:

MoE backend flexible-extract strict-match
Before (nightly 36768d1) Triton 0.9303 +- 0.0070 0.9310 +- 0.0070
After (nightly + PR #55368) AITER (224 -> 256 padded) 0.9269 +- 0.0073 0.9277 +- 0.0072

Server command

# Before (unpatched AITER MoE fails to start at TP8):
# IMAGE=vllm/vllm-openai-rocm:nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89
# AITER_MOE=0
# After:
# IMAGE=vllm/vllm-openai-rocm:nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89  + pr55368
# AITER_MOE=1

docker run -d --init --name k2h-tp8 \
  --network=host --ipc=host --shm-size=64g --device=/dev/kfd --device=/dev/dri \
  --group-add=video --group-add=$(getent group render | cut -d: -f3) \
  --cap-add=SYS_PTRACE --security-opt=seccomp=unconfined \
  -v ${HOME}/.cache/huggingface:/root/.cache/huggingface \
  -e HIP_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
  -e NCCL_MIN_NCHANNELS=112 \
  -e VLLM_ROCM_USE_AITER=1 \
  -e VLLM_ROCM_USE_AITER_MOE=${AITER_MOE} \
  -e AITER_LOG_TUNED_CONFIG=1 \
  -e VLLM_ROCM_USE_AITER_CUSTOM_AR=0 \
  --entrypoint vllm $IMAGE serve IFM/K2-Horizon-375B-A23B \
  --trust-remote-code \
  --tensor-parallel-size=8 \
  --dtype=bfloat16 \
  --gpu-memory-utilization=0.9 \
  --generation-config=auto \
  --reasoning-parser=k2_horizon \
  --enable-auto-tool-choice \
  --tool-call-parser=k2_horizon \
  --no-enable-prefix-caching \
  --max-model-len=8192 \
  --chat-template-content-format=string \
  --attention-backend=ROCM_AITER_UNIFIED_ATTN \
  --max-num-batched-tokens=16384 \
  --port 8000

lm_eval command

lm_eval --model local-completions \
  --model_args model=IFM/K2-Horizon-375B-A23B,base_url=http://localhost:8000/v1/completions,num_concurrent=32,max_retries=5,tokenized_requests=False \
  --tasks gsm8k --num_fewshot 5 --seed 0,1234,1234,1234 --gen_kwargs "temperature=0"

@vllm-agent

vllm-agent commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

CI selector (shadow): 7 test steps (13 jobs) instead of 65 (88 jobs)

Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works.

Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.

steps (jobs) Today's rules Selector Would skip Would add
NVIDIA, CPU and others 65 (88) 7 (13) 59 (78) 1 (3)
AMD mirrors 78 (121) 11 (17) 70 (109) 3 (5)
Selector would run (7)
  • amd-fp8-moe-kernels-mi355
  • basic-models-test-other-cpu
  • kernels-b200 ×3
  • kernels-fusedmoe-layer-test-2-b200s
  • kernels-moe-test ×5
  • kernels-root-misc-test-b200
  • model-executor
Would skip (today's rules run them) (59)
  • ascend-npu-test
  • async-engine-inputs-utils-worker
  • basic-correctness ×2
  • basic-correctness-cpu-offload
  • basic-correctness-cumem
  • basic-correctness-prefetch-offload
  • basic-correctness-sleep-mode
  • basic-models-tests-initialization
  • basic-models-tests-other
  • batch-invariance-b200
  • batch-invariance-h100
  • benchmarks-cli-test
  • cpu-language-generation-and-pooling-model-tests ×3
  • distributed-compile-unit-tests-2xh100
  • entrypoints-integration-api-server ×4
  • entrypoints-integration-api-server-generate
  • entrypoints-integration-api-server-openai-chat_completion
  • entrypoints-integration-api-server-openai-completion
  • entrypoints-integration-llm
  • entrypoints-integration-multimodal
  • entrypoints-integration-pooling
  • entrypoints-integration-responses-api
  • entrypoints-integration-speech_to_text
  • fusion-e2e-quick-h100
  • fusion-e2e-tp2-b200
  • fusion-e2e-tp2-quick-h100
  • kernels-deepgemm-test-h100
  • kernels-fla-ops-test-b200
  • kernels-fusedmoe-layer-test-2-h100s
  • kernels-mhc-test-b200
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • lora-tp-distributed ×4
  • metrics-tracing-2-gpus
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • multi-modal-models-standard-4-other-whisper
  • multi-modal-processor ×4
  • multi-modal-processor-cpu ×4
  • pytorch-compilation-dynamic-shapes
  • pytorch-compilation-passes-unit-tests
  • pytorch-compilation-unit-tests
  • pytorch-compilation-unit-tests-h100
  • pytorch-fullgraph-cudagraph-l4-compatibility
  • pytorch-fullgraph-test
  • quantized-moe-test-b200
  • regression
  • samplers-multimodal-beam-search
  • samplers-test
  • v1-core
  • v1-executor-worker
  • v1-kv-connectors ×4
  • v1-kv-offload
  • v1-logits-oracle
  • v1-metrics-lmeval
  • v1-sample
  • v1-spec-decode
Would add (today's rules do not run them) (1)
  • kernels-b200 ×3 (code map)
AMD mirrors: would skip (70)
  • async-engine-inputs-utils-worker
  • basic-correctness ×2
  • basic-correctness-cpu-offload
  • basic-correctness-cumem
  • basic-correctness-prefetch-offload
  • basic-correctness-sleep-mode
  • basic-models-tests-extra-initialization ×14
  • basic-models-tests-initialization
  • basic-models-tests-other
  • batch-invariance-h100
  • benchmarks-cli-test
  • deepseek-v4-kernel-test-b200
  • distributed-model-tests-2-gpus ×3
  • entrypoints-integration-api-server ×4
  • entrypoints-integration-api-server-generate
  • entrypoints-integration-api-server-openai-chat_completion
  • entrypoints-integration-api-server-openai-completion
  • entrypoints-integration-llm
  • entrypoints-integration-multimodal
  • entrypoints-integration-pooling
  • entrypoints-integration-responses-api
  • entrypoints-integration-speech_to_text
  • fusion-and-compile-unit-tests-2xb200
  • fusion-e2e-quick-h100
  • fusion-e2e-tp2-ar-rms-config-sweep-h100
  • fusion-e2e-tp2-b200
  • fusion-e2e-tp2-quick-h100
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • language-models-tests-extra-standard ×2
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • lora-tp-distributed ×4
  • metrics-tracing-2-gpus
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • multi-modal-models-standard-4-other-whisper
  • multi-modal-processor ×4
  • multi-modal-processor-cpu ×4
  • openai-api-correctness
  • pipeline-context-parallelism-4-gpus
  • platform-tests
  • pytorch-compilation-dynamic-shapes
  • pytorch-compilation-passes-unit-tests
  • pytorch-compilation-unit-tests
  • pytorch-compilation-unit-tests-h100
  • pytorch-fullgraph-test
  • quantization ×4
  • quantized-moe-test-b200
  • regression
  • samplers-multimodal-beam-search
  • samplers-test
  • spec-decode-draft-model ×4
  • spec-decode-eagle-1-deepseek-qwen
  • spec-decode-eagle-2-llama3-qwen-vl-other
  • spec-decode-mtp-deepseek-mimo
  • spec-decode-mtp-gemma4
  • spec-decode-mtp-qwen3-5
  • spec-decode-ngram-suffix
  • spec-decode-speculators
  • v1-core
  • v1-executor-worker
  • v1-kv-connectors ×4
  • v1-kv-offload
  • v1-logits-oracle
  • v1-metrics-lmeval
  • v1-sample
  • v1-spec-decode
AMD mirrors: would add (3)
  • deepseek-v4-kernel-test-h100 (code map)
  • gemm-rs-ar-2xb200 (code map)
  • kernels-core-operation-test ×3 (code map)

CI results (2026-09-30 15:52 UTC)

166 passed, 0 failed, 0 pending.

No failures to judge.

4 changed files · base 2df122e65e · head 4e1182a3a6 · Python record: not used (/tmp/ci-infra-selector/buildkite/ci_selector/coverage-data/table.json.gz is table version 5, expected 7; re-merge it from the raw recordings) · kernel record: table 866fa13 (build 92059), map 866fa13 · not counted: 10 build steps, 5 A100 steps the generator no longer emits, 2 optional steps the selector would also run

@shen-shanshan
shen-shanshan merged commit 4e1182a into vllm-project:main Sep 30, 2026
180 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready ONLY add when PR is ready to merge/full CI is needed rocm Related to AMD ROCm verified Run pre-commit for new contributors without triggering other tests

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

8 participants